Papers with subjective metrics
Versatile Framework for Song Generation with Prompt-based Control (2025.findings-emnlp)
Copied to clipboard
Yu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu, Ruiqi Li, Jingyu Lu, Rongjie Huang, Ruiyuan Zhang, Zhiqing Hong, Ziyue Jiang, Zhou Zhao
| Challenge: | Existing methods for song generation fail to generate vocals with prompt-based control and proper alignment. |
| Approach: | VersBand is a multi-task song generation framework for synthesizing high-quality songs with prompt-based control. |
| Outcome: | Experimental results show that VersBand performs better than baseline models across multiple song generation tasks. |
DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility (2025.findings-naacl)
Copied to clipboard
| Challenge: | Video-to-speech (V2S) synthesis requires acoustic hints to accurately reconstruct both speech content and speaker characteristics from video clips alone. |
| Approach: | They propose a video-to-speech (V2S) model that predicts Mel-spectrograms directly from video frames. |
| Outcome: | The proposed model outperforms existing models in acoustic intelligibility and preserves speaker-specific characteristics. |
In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis (2026.acl-long)
Copied to clipboard
| Challenge: | citation counts are a shallow view that fails to capture how a paper has influenced subsequent work. |
| Approach: | They propose a task to generate nuanced, expressive, and time-aware impact summaries . they analyze fine-grained confirmatory and correction citation intents to generate summary . |
| Outcome: | The proposed task shows moderate to strong human correlation on subjective metrics such as insightfulness. |
AlignSTS: Speech-to-Singing Conversion via Cross-Modal Alignment (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to speech-to-singing voice conversion are difficult to learn in text-free situations. |
| Approach: | They propose an STS model which views speech variance as different modalities . it uses a novel rhythm adaptor to predict the target rhythm representation . they also use the predicted rhythm representation to re-align the content . |
| Outcome: | The proposed model achieves superior performance in terms of objective and subjective metrics. |
Learning the Beauty in Songs: Neural Singing Voice Beautifier (2022.acl-long)
Copied to clipboard
| Challenge: | Existing techniques for pitch correction are limited to intonation but ignore the overall aesthetic quality. |
| Approach: | They propose a novel time-warping approach for pitch correction to synchronize the amateur recording with the template pitch curve. |
| Outcome: | The proposed model improves intonation and vocal tone while keeping content and vocal timbre. |
UniCoM: A Universal Code-Switching Speech Generator (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Code-switching (CS) is a common phenomenon in real-world conversations and poses significant challenges for multilingual speech technology. |
| Approach: | They propose a pipeline for generating high-quality, natural CS samples without altering sentence semantics. |
| Outcome: | The proposed pipeline generates high-quality, natural CS samples without altering sentence semantics without alteration of sentence semantic. |
FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for speech editing still suffer from over-smoothing problem and lack of robustness due to stutter. |
| Approach: | They propose a stutter-oriented automatic speech editing model that incorporates sutter information into the hidden sequence. |
| Outcome: | The proposed model achieves state-of-the-art performance on a speech recording dataset . it can improve fluency of stuttering speech in terms of objective and subjective metrics. |
Mind Reader: Latent User Demand-Guided Content Optimization for Generative Search Engine (2026.acl-long)
Copied to clipboard
Tong Chen, JiaWei Guo, Yuxi Li, Baiming Chen, Houxing Ren, Zhang Zhiwei, Yunxiang Zhang, Hanyang Xia, Kun Liang, Zhaoran Fan
| Challenge: | Generative Search Engines (GSEs) have reshaped information retrieval and Generating Engine Optimization (GEO) emerges to improve the content visibility in GSEs’ responses. |
| Approach: | They propose a method to optimize content to cover latent semantic information of GSEs by decomposing query into diverse perspectives and capturing underlying semantic information. |
| Outcome: | The proposed method outperforms baselines and effectively improves content visibility (with up to 2.44x objective metrics and 1.23x subjective metrics on average). |
Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music (2026.acl-long)
Copied to clipboard
| Challenge: | Existing symbolic music generation models represent musical notes as a sequence of attribute tokens with fixed unidirectional dependencies. |
| Approach: | They propose a symbolic music generation framework that adopts a autoregressive and a discrete diffusion architectures for note attributes. |
| Outcome: | The proposed framework improves state-of-the-art models across objective and subjective metrics. |